Repository navigation
Conversation
|
Semgrep found 12
🟡 Medium severity issue identified in your code: The method identified is susceptible to injection. The input should be validated and properly If this is a critical or high severity finding, please also link this issue in the #security channel in Slack. |
|
Claude finished @jcastro-dotcms's task in 1m 52s —— View job PR Review — S3 binary backfill, starter export/import, integrity repairI reviewed the diff against New Issues
Existing
Notes
Only the High finding is blocking; it's a one-line guard already written in the parallel path. · |
05b261e to
4ae1a5c
Compare
|
Semgrep found 4
🟡 Medium severity issue identified in your code: The method identified is susceptible to injection. The input should be validated and properly If this is a critical or high severity finding, please also link this issue in the #security channel in Slack. |
|
test |
|
Pull Request Unsafe to Rollback!!!
AI rollback-safety analysis for commit range |
4ae1a5c to
86992b7
Compare
86992b7 to
4f778fe
Compare
165a660 to
5824a30
Compare
5824a30 to
ea2656b
Compare
4f778fe to
f552bdb
Compare
ea2656b to
f874dda
Compare
f552bdb to
13879e9
Compare
13879e9 to
6771ed3
Compare
|
…grity repair Fourth slice of the S3 asset storage work. With FEATURE_FLAG_S3_ASSET_STORAGE on, administrators can migrate existing binaries to S3 through the checkpointed binaryAssetBackfill job, starter export restores originals from S3 and starter import publishes binaries before its database commit, and the file-asset integrity repair copies stored binaries with the corrected content. The backfill processor does not register while the flag is off, and flag-off behavior matches main.
… and fail exports unreadably These fixes apply only with FEATURE_FLAG_S3_ASSET_STORAGE on; flag-off export, import and error handling are unchanged. The backfill batch query and the starter export's content scan now put the page size in the SQL through DotConnect.setSQL(sql, limit). DotConnect.setMaxRows only trimmed rows after the driver had fetched the whole remaining table. The export now pages on inode alone and loads each accepted row's JSON separately, so a page never holds the table's JSON. Backfill and starter import share one policy for data problems a retry cannot fix: a referenced binary or metadata record that exists neither locally nor in S3, a row whose JSON cannot be parsed, and a legacy row whose content cannot be found are logged and skipped, and the rest of the row is still copied. Import logs a summary count, so starters exported without assets, with maxSize or with oldAssets=false import again, including on first boot. S3 and database errors still fail the batch or the import. The backfill job continues past skipped inodes and reports skippedCount and the first 1000 skippedInodes in its checkpoint and result; storage failures still fail the batch, and the job error now carries the failing inode. The checkpoint compare-and-set is unchanged. The job sends a heartbeat after every inode. A failed flag-on export no longer finishes the ZipOutputStream, so the client receives a truncated archive without a central directory instead of a valid partial backup. Export takes the cache lease per binary rather than for the whole export, skips a binary whose stored metadata exceeds maxSize before restoring it, and holds one lease only around the local folder walk, which reads already-local files. The file-asset integrity repair registers rollback listeners that delete the revision objects and metadata it uploaded if the repair transaction rolls back; a failed deletion is logged. The behavior doc describes these changes and states that a starter exported with the flag on cannot be fully restored into a flag-off instance.
6771ed3 to
ed86935
Compare
3523198 to
236a0d2
Compare
| if (!filter.accept(new File(ConfigUtils.getAssetPath(), inodePath))) { | ||
| continue; | ||
| } | ||
| final Contentlet content = new Contentlet(APILocator.getContentletAPI().find(inode, APILocator.systemUser(), false)); |
There was a problem hiding this comment.
🔴 [P1] ExportStarterUtil.java:544 guard null find before copy
Current code:
final Contentlet content = new Contentlet(APILocator.getContentletAPI().find(inode, APILocator.systemUser(), false));Problem: find() can return null, copy then NPEs.
Fix:
final Contentlet found = APILocator.getContentletAPI().find(inode, APILocator.systemUser(), false);
if (found == null) {
Logger.warn(this, "Skipping unreadable content " + inode);
continue;
}
final Contentlet content = new Contentlet(found);| binaries += copyInode(inode, skipped); | ||
| cursor = inode; | ||
| afterEachInode.run(); | ||
| } |
There was a problem hiding this comment.
🟡 [P2] BinaryAssetBackfill.java:74 concurrent deletes can end the scan prematurely with complete=true
Current code:
return new Result(cursor, binaries, rows.size() < limit, List.copyOf(skipped));Problem: The batch is marked "complete" whenever a page returns fewer rows than limit, but a concurrent content deletion between page fetches can shrink a page below limit while later inodes still exist.
Fix:
final boolean lastPage = rows.size() < limit && new DotConnect()
.setSQL("select 1 from contentlet where inode > ?", 1)
.addParam(cursor).loadObjectResults().isEmpty();
return new Result(cursor, binaries, lastPage, List.copyOf(skipped));Why it matters: the admin-triggered BinaryAssetBackfillProcessor job runs against live traffic. If a page comes up short because rows were deleted, the job persists complete=true in its checkpoint and stops, so remaining content versions are never verified in S3, yet the job reports success. Starter import (runAll) is unaffected because it runs inside the import transaction.
|
dotbot code review:
No new actionable bugs were found in the current changes, but 2 prior unresolved dotbot findings still apply, so the patch remains incorrect. Tip: comment with "/dotbot address comments" to attempt automated fixes for unresolved review threads. reviewed by dotbot · meta/muse-spark-1.3 · medium |
|
dotbot code review:
No new actionable bugs were found in the current changes, but 2 prior unresolved dotbot findings still apply, so the patch remains incorrect. Tip: comment with "/dotbot address comments" to attempt automated fixes for unresolved review threads. reviewed by dotbot · ~z-ai/glm-latest · medium |
Refs #37868
Proposed Changes
POST /api/v1/jobs/binaryAssetBackfill, administrators only, checked at queue and run time). Copies existing originals, metadata and recognized renditions to S3 in verified batches, advancing a persisted cursor only after a fully verified batch. Resumable and cancellable, never deletes local sources, and rejected while the flag is off. A row whose binary or metadata exists nowhere, or whose JSON cannot be read, is skipped and reported in the job result (skippedCountand up to 1000skippedInodes); a storage error still fails the batch and names the inode.maxSizeor witholdAssets=falsestill import; S3 and database errors still fail the import, and so does an error importing rules, so a starter is never left partly imported. Importing over a populated database with the flag on performs a full replacement with foreign keys enforced.Behavior with the flag off
Unchanged from main; the backfill processor does not register.
Review fixes
The last commit on this branch (
fix(storage): bound recovery scans, tolerate missing starter binaries and fail exports unreadably) addresses a full review of this PR. All of it is flag-on only:DotConnect.setMaxRowsonly trims rows after they are fetched, so each batch used to read the whole remaining contentlet table, and export loaded every row's JSON into memory. Export pages oninodeand loads each row's JSON on its own.Rollback safety
Backfill converts raw S3 objects into SHA-256 references in place, which an older release cannot read. Same rule as 2: once the flag is on, treat it as forward-only.
Checklist
BinaryAssetBackfillTestandExportStarterFailureTest. Integration on this branch before the review fixes, flag off: 34 run, 0 failures, including the existingContentFileAssetIntegrityCheckerTestandESMappingAPINumericFieldTest; flag on: 23 run, 0 failures.BinaryAssetStarterRestoreTest(registered inJunit5Suite1, opt-in) passed all phases against one MinIO bucket before the review fixes: export, fresh-database restore from the exported ZIP, and replacement of a populated database. After the review fixes, the integration suites were run on the top of the stack (feat(storage): serve renditions, compiled CSS and templates through S3 asset storage #37776); the starter phases were not rerun.